AITopics

Neural Information Processing SystemsApr-26-2026, 17:39:52 GMT

Bandits

algorithm, artificial intelligence, machine learning, (18 more...)

Technology: Information Technology > Artificial Intelligence > Machine Learning (1.00)

Bandit Phase Retrieval

Neural Information Processing SystemsApr-26-2026, 17:39:48 GMT

We study a bandit version of phase retrieval where the learner chooses actions (At)nt=1 in the d-dimensional unit ball and the expected reward is hAt,?i2 with? 2 Rd an unknown parameter vector. We prove an upper bound on the minimax cumulative regret in this problem of (d p n), which matches known lower bounds up to logarithmic factors and improves on the best known upper bound by a factor of p d. We also show that the minimax simple regret is (d/ p n) and that this is only achievable by an adaptive algorithm. Our analysis shows that an apparently convincing heuristic for guessing lower bounds can be misleading and that uniform bounds on the information ratio for information-directed sampling [Russo and Van Roy, 2014] are not sufficient for optimal regret.

algorithm, artificial intelligence, machine learning, (18 more...)

Technology: Information Technology > Artificial Intelligence > Machine Learning (1.00)

Bandit Phase Retrieval

Neural Information Processing SystemsFeb-10-2026, 05:24:47 GMT

Roy, 2014] are not sufficient for optimal regret.

algorithm, artificial intelligence, machine learning, (18 more...)

Country:

North America > United States (0.04)
Europe > United Kingdom > England > Cambridgeshire > Cambridge (0.04)

Technology: Information Technology > Artificial Intelligence > Machine Learning (1.00)

Bandit Phase Retrieval

Neural Information Processing SystemsFeb-10-2026, 05:24:44 GMT

Bandits

algorithm, bandit, retrieval, (16 more...)

Country:

North America > United States (0.04)
Europe > United Kingdom > England > Cambridgeshire > Cambridge (0.04)

Technology: Information Technology > Artificial Intelligence > Machine Learning (1.00)

Neural Information Processing SystemsDec-24-2025, 14:11:30 GMT

Bandit Phase Retrieval

We study a bandit version of phase retrieval where the learner chooses actions $(A_t)_{t=1}^n$ in the $d$-dimensional unit ball and the expected reward is $\langle A_t, \theta_\star \rangle^2$ with $\theta_\star \in \mathbb R^d$ an unknown parameter vector. We prove an upper bound on the minimax cumulative regret in this problem of $\smash{\tilde \Theta(d \sqrt{n})}$, which matches known lower bounds up to logarithmic factors and improves on the best known upper bound by a factor of $\smash{\sqrt{d}}$. We also show that the minimax simple regret is $\smash{\tilde \Theta(d / \sqrt{n})}$ and that this is only achievable by an adaptive algorithm. Our analysis shows that an apparently convincing heuristic for guessing lower bounds can be misleading and that uniform bounds on the information ratio for information-directed sampling (Russo and Van Roy, 2014) are not sufficient for optimal regret.

bandit phase retrieval, electronic proceedings, name change, (4 more...)

Technology: Information Technology > Artificial Intelligence > Machine Learning (0.44)

Neural Information Processing SystemsFeb-9-2025, 12:39:41 GMT

Bandit Phase Retrieval

We study a bandit version of phase retrieval where the learner chooses actions (A_t)_{t 1} n in the d -dimensional unit ball and the expected reward is \langle A_t, \theta_\star \rangle 2 with \theta_\star \in \mathbb R d an unknown parameter vector. We prove an upper bound on the minimax cumulative regret in this problem of \smash{\tilde \Theta(d \sqrt{n})}, which matches known lower bounds up to logarithmic factors and improves on the best known upper bound by a factor of \smash{\sqrt{d}} . We also show that the minimax simple regret is \smash{\tilde \Theta(d / \sqrt{n})} and that this is only achievable by an adaptive algorithm. Our analysis shows that an apparently convincing heuristic for guessing lower bounds can be misleading and that uniform bounds on the information ratio for information-directed sampling (Russo and Van Roy, 2014) are not sufficient for optimal regret.

bandit phase retrieval, smash, tilde theta, (1 more...)

Technology: Information Technology > Artificial Intelligence (0.71)

Lattimore, Tor, Hao, Botao

Bandit Phase Retrieval

arXiv.org Machine LearningJun-4-2021

We study a bandit version of phase retrieval where the learner chooses actions $(A_t)_{t=1}^n$ in the $d$-dimensional unit ball and the expected reward is $\langle A_t, \theta_\star\rangle^2$ where $\theta_\star \in \mathbb R^d$ is an unknown parameter vector. We prove that the minimax cumulative regret in this problem is $\smash{\tilde \Theta(d \sqrt{n})}$, which improves on the best known bounds by a factor of $\smash{\sqrt{d}}$. We also show that the minimax simple regret is $\smash{\tilde \Theta(d / \sqrt{n})}$ and that this is only achievable by an adaptive algorithm. Our analysis shows that an apparently convincing heuristic for guessing lower bounds can be misleading and that uniform bounds on the information ratio for information-directed sampling are not sufficient for optimal regret.

algorithm, bandit, retrieval, (16 more...)

arXiv.org Machine Learning

2106.0166

Country:

North America > United States (0.04)
Europe > United Kingdom > England > Cambridgeshire > Cambridge (0.04)

Genre: Research Report (0.82)

Technology: Information Technology > Artificial Intelligence > Machine Learning (1.00)